npj Precision Oncology
○ Springer Science and Business Media LLC
Preprints posted in the last 30 days, ranked by how well they match npj Precision Oncology's content profile, based on 53 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.
Schuerch, M.; Geisberg, J.; Flower, C. T.; Bektas, A. B.; McDonald, T. O.; Mishra, S.; Graser, C.; Altreuter, J.; Ananda, G.; Boland, G.; Liu, D.; kehl, K. L.; Michor, F.
Show abstract
Progress in precision oncology, including biomarker discovery and individualized treatment selection, is limited by the complexity of clinico-genomic data and the scarcity of large multimodal patient cohorts. Here, we introduce PanoraOnc, a pan-cancer artificial intelligence (AI) model pretrained on real-world clinical, genomic, and imaging data from 84,131 patients spanning 66 cancer types. PanoraOnc enables transferable treatment outcome prediction through pan-cancer pretraining and generalizes to unseen cohorts across cancer types, institutions, and therapeutic settings. Evaluation and fine-tuning were performed on cohorts comprising diverse modalities, including clinical features, targeted gene panels, immunofluorescence imaging, whole-exome sequencing, and transcriptomic profiles. Across these settings, PanoraOnc consistently outperforms statistical, machine-learning, survival, and AI baselines, with the largest improvements observed in zero- and few-shot scenarios, demonstrating that large-scale clinico-genomic pretraining enables robust and generalizable outcome predictions across previously unseen conditions. In addition, PanoraOnc supports biomarker discovery through explainable AI, revealing both established and underappreciated features, including tumor-infiltrating clonal hematopoiesis, oncogenic signaling pathways, and DNA damage response mechanisms in immunotherapy-treated melanoma and non-small cell lung cancer. Furthermore, PanoraOnc enables the identification of patient subgroups potentially benefitting from alternative treatments by estimating personalized treatment outcomes across therapeutic scenarios. These findings establish pan-cancer multimodal pretraining as a scalable paradigm for AI-assisted discovery in precision oncology.
Seyedshahi, F. A.; Damiola, F.; Sequeiros, R.; Forest, F.; Scherpereel, A.; Yuan, K.; Lantuejoul, S.; Le Quesne, J.
Show abstract
1Accurate subtype diagnosis is essential for guiding therapy and predicting patient outcome in malignant mesothelioma. Most computational pathology models are trained on large tissue images from resection specimens, which maximises information for training but limits model relevance in real-world diagnostic settings where small biopsies are the most usual tissue source. In this work, we assembled a large multicentre cohort of HES- and HPS-stained mesothelioma biopsy slides. We used a self-supervised learning model to evaluate the associations of biopsy-driven morphology patterns with histological subtype, molecular markers, and survival. The discovered histomorphology patterns captured a continuum of tissue phenotypes spanning epithelioid, sarcomatoid, and non-tumour morphologies. Also, patient-level HPC representations achieved excellent performance for distinguishing epithelioid from non-epithelioid mesothelioma (AUC = 0.94) and demonstrated predictive value for immunohistochemistry (IHC) markers. Additionally, HPC-derived features alone achieved performance comparable to established clinical and molecular variables (C-index = 0.65), while integration of HPCs with clinical and marker information improved performance to a C-index of 0.69. Several HPCs were significantly associated with favourable or adverse prognosis and reflected known subtype-specific biological patterns. In conclusion, self-supervised learning can discover interpretable histomorphological phenotypes directly from routine mesothelioma biopsies without further training. These AI-derived phenotypes capture clinically and biologically relevant information, linking tissue architecture to molecular characteristics, histological subtypes, and patient outcomes. The proposed framework provides a thorough evaluation of real-world biopsy data using a pre-trained model, without the need for computationally intensive retraining, and addresses the question of whether SSL-based AI can be deployed out of the box in clinical settings.
Yang, Y.; Vasudevaraja, V.; Serrano, J.; Mohamed, H.; Kelly, S.; Jour, G.; Gindin, T.; Park, K.; Jones, D.; Feng, X.; Pinnell, J.; Mclennan, S.; Tin, M. Y.; Tsirigos, A.; Snuderl, M.; Wrzeszczynski, K. O.
Show abstract
Next-generation sequencing (NGS) for the detection of somatic variants has become the method of choice in a variety of molecular oncology fields and in the clinic. Its use ranges from sequencing entire tumor genomes and transcriptomes to targeted clinical diagnostic gene panels. The NYU Langone Genome PACT (Profiling of Actionable Cancer Targets, LG-PACT) assay is a qualitative in vitro diagnostic test that uses targeted next generation sequencing (NGS) of formalin-fixed paraffin-embedded (FFPE) tumor tissue matched with normal specimens from patients to detect gene alterations in a targeted panel covering 606 genes and the TERT promoter. Indications for testing are cancer (solid tumors and hematological malignancies) where a mutational profile from multiple genes would be informative for disease stratification, prognosis, or treatment options including targeted therapies and eligibility for clinical trials. The test is intended to provide information on somatic mutations including point mutations, small insertions/deletions (indels), and copy number aberrations for diagnostic and treatment decisions. LG-PACT is a United States Food and Drug Administration (FDA) cleared diagnostic test (510K: K202304). The clinical interpretation of sequencing data of molecular tumor markers from NGS encompasses automated variant calling tools with human interpretation. This final mostly manual review of data step is intensive, involving highly trained scientists, encompassing literature review, interpretation and clinical tier classification by pathologists, who then provide a complete molecular diagnostic report to the treating oncologists. We provide analysis of 1339 clinical genomic profiles from 31 different cancers and their subtypes, comprising of central nervous system (CNS) 792 (59%) cases (incl. meningioma, glioma and glioblastoma), with 267 (20%) cases predominantly of lung, pancreatic and colorectal and 280 of others (21%). Here, we present the technical challenges of validating an NGS oncological diagnostic targeted assay for clinical grade accuracy and sensitivity for patient care. We show how copy number alterations provide a more comprehensive description of the tumors genomic profile. We then outline the utility of targeted panel sequencing based on certified pathologist selection of reportable variants for our current patient cohort. Where analysis of variant detection has led to 49.4% (661/1339) of our clinical tumor samples containing mutations in known therapy targeted genes, 35.6% (477/1339) with mutation detected in other genes, and 15% (201/1339) cases being negative.
Imran, A.; Rahat Hossain, K. M.; Islam, S. M. R.; Rahman, M. S.
Show abstract
Triple-Negative Breast Cancer (TNBC) is characterized by high heterogeneity, poor prognosis, and limited targeted treatment options. Bridging the gap between molecular alterations and histopathological morphology remains a major challenge in precision oncology. We propose an interpretable, multi-modal framework that integrates histopathological image analysis with multi-omics profiling (somatic mutations, DNA methylation, copy number alterations), leveraging U-Net-based nuclei segmentation, vision-language models (BLIP), biomedical language models (BioGPT), and explainable AI (SHAP, LIME). Our framework achieves strong predictive performance (AUC = 0.989) and provides transparent, biologically grounded interpretations by integrating morphological features with genomically prioritized biomarkers. Cross-modal analysis confirms established TNBC drivers and generates novel, testable hypotheses associating specific epigenetic alterations with distinct morphological phenotypes. While causal validation requires future wet-lab experiments, our framework accelerates hypothesis-driven biomarker discovery by integrating complementary data modalities with language-based reasoning, providing a transparent foundation for hypothesis generation and clinical translation.
Gao, Y.; Yu, S.; Xia, Y.; Chen, S.; Xia, S.; An, R.; Zeng, J.; Zhao, F.; Ma, Y.; Wang, Y.; Xie, X.; Zhang, J.
Show abstract
Prognostic models in oncology are developed one cancer at a time, from that cancer's own labelled outcomes, and fail where prognostic information is scarcest. Rare cancers account for roughly a fifth of diagnoses and most paediatric malignancies, yet seldom supply enough events for a reliable time-to-event model. We therefore asked whether a representation learned without outcome labels can supply what those cohorts cannot. A Transformer encoder was pretrained by masked field-value modelling on 9425135 tumour records from the SEER 17 registries, diagnosed in 2000 to 2023. Only diagnosis-time fields passing a fail-closed coding-verification gate were admitted, and each record was emitted as an era-specific and a harmonised view, keeping two decades of recoding auditable. The encoder was then frozen and read by a linear Cox head for overall survival. Nine rare cancers were removed from the pretraining corpus entirely, each requiring an independent pretraining run. On a sealed test partition, all nine exceeded an architecture-identical random frozen encoder in Harrell concordance by +0.0034 to +0.0368, every lower confidence limit above zero. At 256 labelled patients, all 67 cancers favoured the pretrained representation over budget-matched Cox regression, median difference +0.0283. The advantage was bounded: given the entire training set, Cox regression was favoured in seven of nine rare cancers. The encoder did not outperform a field-frequency baseline on its own objective, so upstream reconstruction did not predict downstream transfer. Outcome-agnostic registry pretraining carries prognostic signal into cancers it has never seen, and is most useful where labels are fewest, without establishing clinical utility.
Schulze, F.; Loeffler, C.; Radoynova, M.; Winter, S.; Roellig, C.; Sockel, K.; Kroschinsky, F.; Bornhaeuser, M.; Middeke, J. M.; Kather, J. N.; Eckardt, J.-N.; Ghaffari Laleh, N.
Show abstract
Hematologic diagnostics and especially cytomorphologic assessment are time-intensive and require high levels of expertise. Vision Language Models (VLM) show promise in medical image analysis in radiology and histopathology, while an evaluation on detecting acute myeloid leukemia (AML) is lacking. Our goal was to evaluate three Vision Language Models regarding their diagnostic accuracy and safety in clinical decision support in detecting AML from digitized bone marrow smears (BMS). Whole slide images were obtained from bone marrow smears of 50 AML patients and 50 bone marrow donors. Ten representative fields of view per sample were extracted manually. Three VLMs were used, two of which are considered generalist models (Qwen3.5-397B-A17B-FP8, GLM-4.6V-FP8), while the other one is a medically adapted model (Medgemma-27b-it). All models performed zero-shot analysis using two prompting strategies: First, a context-rich prompt requesting reporting of WHO/FAB diagnostic criteria in a structured manner, and secondly a minimal prompt without specific hematologic context. Overall diagnostic accuracy was poor for all models as they exhibited the overwhelming tendency to classify most samples as leukemic: With context-rich prompts, GLM4.6 identified 90% of leukemic samples while also labeling 92% of bone marrow donors as AML. The medical specialist model MedGemma-27b showed similar failure, misclassifying 86% of healthy donors and correctly detecting AML in only 66% of cases. Qwen3.5 performed best under detailed prompting, achieving a specificity of 0.26 and accuracy of 0.51. Accuracy of all models improved with context-free prompts (accuracies range 0.47-0.79), yet they still lacked the ability to correctly distinguish between leukemia and healthy bone marrow. Qwen3.5 was the only model to maintain meaningful specificity (0.64) and correctly identified 94% of AML, yielding an overall accuracy of 0.79. Morphologic feature-level agreement with human expert reports was poor across all models, indicating poor recognition of cell-level morphologies. This failure is likely driven by the fact that pathology imaging archives are vastly scraped during model training while hematological samples are not as widely available and therefore, hematology is an out-of-bounds use-case for these models, rendering them currently unsuitable for clinical decision support in hematology.
Peralta Viteri, C.; Harnischfeger, N.; Szabo, L.; Hartmann, S.; Kretzschmar, K.
Show abstract
Precision oncology seeks to match each tumor with the most effective anti-cancer therapy. Advances in pharmacogenomics and machine learning enabled drug response prediction models with strong performance in cancer cell lines. Nonetheless, patient-centric evaluation of drug prioritization and systematic assessment of model generalization in patient-derived systems across cancer types remain largely absent. Here we introduce a translational framework combining patient-centric benchmarking with a pan-cancer pharmacogenomic atlas of patient-derived organoids, together with NELLY, a deep learning model integrating transcriptomic and chemical information to predict drug response and prioritize therapies. NELLY outperformed existing methods for patient-specific drug prioritization across cancer cell lines and patient-derived organoids, including under out-of-distribution evaluation. Its dynamic weighting mechanism provided patient-specific gene attributions, offering a route to connect predicted drug response to molecular programs associated with drug resistance. Our results support NELLY as a promising framework for translationally relevant and interpretable drug response prediction in precision oncology.
Zhu, M.; Li, A.; Safa, I.; Galera, P.; Hazoglou, M.; Vanderbilt, C.; Kamali, A.; Goldgof, G.; Veeraraghavan, H.; Jiang, J.; Ardon, O.; Geneslaw, L.; Dogan, A.
Show abstract
Pathologic diagnoses of hematopoietic diseases require immunohistochemistry (IHC) stains selected by pathologists upon preview of H&E-stained slides. This multi-step workflow can delay diagnostic turnaround time by days. Hence, we developed the Hematopathology Automatic Triaging System (HATS), which automates IHC panel ordering directly from H&E whole-slide images using pretrained pathology foundation model representations combined with attention-based multiple-instance learning. After the most comprehensive evaluation of pathology foundation models for hematologic malignancy classification to date, encompassing seven publicly available models, we trained HATS on 4,996 whole-slide images from 1,607 patients spanning the ten most common lymphoma diagnostic categories. HATS achieves 84% case-level subtype classification accuracy (0.962 ROC-AUC), translating to 92% IHC panel ordering accuracy. In a blinded reader study, HATS outperforms practicing pathologists at predicting lymphoma subtypes from morphology alone (85% vs 65%). In an independent real-world validation of 230 clinical cases, after directing 7 cases with scant tissue for manual review, HATS-ordered IHC panels were sufficient for diagnosis in 72.6% of cases. By automating the triaging step while preserving full pathologist oversight, HATS offers a safe and practical entry point for clinical AI adoption in pathology.
Chauhan, S.; Jones, K.; Krajbich, V. A.; Smith, B.; McCallister, C.; Bui, T.; Smith, R.; Woltjer, R. L.; Wangsiricharoen, S.; Ramsay, D.; Davare, M. A.
Show abstract
TFCP2-rearranged rhabdomyosarcoma is an exceptionally rare and highly aggressive malignancy driven by TFCP2 gene fusions and associated with a dismal clinical prognosis. Because standardized treatment regimens are lacking, developing representative preclinical models is critical for identifying effective therapies. Here, we present a case of a 29-year-old male with rapidly progressive, metastatic pelvic intraosseous rhabdomyosarcoma (iRMS) harboring a FUS::TFCP2 fusion and anaplastic lymphoma kinase (ALK) overexpression. To evaluate therapeutic vulnerabilities, we established a patient-derived xenograft (PDX) model that faithfully recapitulated the histologic, immunohistochemical, and molecular hallmarks of the primary tumor. High-throughput in vitro pharmacological screening of PDX-derived cells demonstrated notable resistance to standard cytotoxic chemotherapies and revealed a paradoxical and selective sensitivity profile across ALK inhibitors. The PDX-derived cells were susceptible to crizotinib, brigatinib, and ceritinib, yet resistant to the more selective second- and third-generation inhibitors alectinib and lorlatinib. Notably, next-generation ROS1/pan-TRK inhibitors (entrectinib, repotrectinib, and taletrectinib) demonstrated superior efficacy compared to the fourth-generation ALK inhibitor NVL-655. Our findings establish a validated preclinical PDX model for FUS::TFCP2 iRMS and suggest that multi-targeted tyrosine kinase inhibition may offer a more viable therapeutic strategy than narrow-spectrum ALK targeting or conventional chemotherapy.
Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.
Show abstract
Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.
Naucke, C.; Rodland, G. E.; Eek Mariampillai, A.; Hauge, S.; Steive, L. H.; Bjerke, I. A.; Lindbergsengen, L.; Grosvik, A. S. G.; Siggerud, V.; Kongsrud, K.; Savu, D. I.; Stokke, T.; Syljuasen, R. G.
Show abstract
Radiotherapy induces cytotoxic DNA damage, but activation of DNA repair pathways and cell-cycle checkpoints can limit therapeutic efficacy. Here, we developed a high-throughput, flow cytometry-based screening platform to identify compounds that inhibit radiation-induced DNA repair and checkpoint activation. Reh leukemia and A549 lung cancer cells were irradiated and screened against up to 700 bioactive compounds, with DNA damage persistence quantified by {gamma}H2AX levels across independent screens. Cell barcoding using Pacific Blue staining was incorporated to enable highly accurate quantification of {gamma}H2AX across treatment conditions. The platform yielded robust and reproducible results and supported multiparametric analysis, including assessment of G2 checkpoint activation by phospho-histone H3. Largely overlapping candidate radiosensitizers were identified in both cell lines, including the multi-kinase inhibitor 5-iodotubercidin and the PI3K/mTOR inhibitor omipalisib. Validation studies in lung cancer and glioblastoma models confirmed screen performance. Mechanistically, omipalisib reduced phosphorylation of the non-homologous end-joining protein DNA-PK, consistent with impaired double-strand break repair. Both compounds enhanced radiosensitivity in clonogenic survival assays. Notably, 5-iodotubercidin increased radiosensitivity in glioblastoma cells despite previous reports of radioprotective effects in normal brain tissue. Together, these findings establish a robust barcoded screening approach for identifying radiosensitizers that target DNA damage repair and checkpoint responses.
Majumder, B. P.; Linak, J. A.; Adamson, R.; Aguilera, R. L.; Agarwal, D.; Reitz, Z.; Loiselle, S.; Devarakonda, S.; Clark, P.; Paulson, K. G.; Stanton, S.
Show abstract
In large data sets discovery is often limited to pre-conceived hypotheses and data fishing. Here we tested whether systematic exploration of AI generated hypotheses could uncover clinically meaningful signals in extensively studied data. We deployed AutoDiscovery, a newly launched large language model (LLM) framework designed to search for hypotheses based on surprisal and systematically interrogate complex datasets, on The Cancer Genome Atlas breast cancer cohort. The system did not identify clinically meaningful novel findings without human input. However, a seeded warm-start run with minimal text input from an oncologist revealed multiple interesting and surprising hypotheses. Among these was that a robust immune signature was present across all subtypes of invasive lobular carcinoma (ILC) that exceeded invasive ductal carcinoma (IDC). This observation was independently validated in independent cohorts and confirmed by high-sensitivity multi-immunofluorescence tumor tissue analyses. These results suggest immunotherapy approaches should be tested in ILC including early stage ER+HER2- ILC; these patients are currently excluded from large neoadjuvant immunotherapy trials. They further demonstrate that surprisal-based hypothesis generation frameworks can extract previously unappreciated patterns from deeply interrogated cancer datasets and imply that disease domain experts working with LLMs can derive more meaningful insights from complex data than either could achieve alone.
Feng, B.-J.; Fatema, K.; Nix, D. A.; Atkinson, A.; Caparas, C.; Stubben, C. J.; Lum, D. H.; Parnell, T. J.; Carroll, C.; Grass, G. D.; Graham, L.; Singer, E. A.; Nepple, K. G.; Manojlovic, Z.; Kauffman, E.; King, J. M.; Ghodoussipour, S.; Hensley, P.; Viscuse, P. V.; Ayanambakkam, A.; Churchman, M. L.; Swami, U.; Agarwal, N.; Cairns, B.; Gupta, S.
Show abstract
PurposeSWI/SNF (BAF) chromatin remodeling complex alterations are common in urothelial carcinoma, yet no biomarker-directed therapeutic strategies have been established for this population. We investigated whether BAF alterations delineate a biologically distinct, therapeutically actionable urothelial carcinoma subtype. Experimental DesignWe performed integrative genomic and transcriptomic analyses of 792 urothelial carcinoma tumors from the Oncology Research Information Exchange Network (ORIEN) and validated findings in the TCGA-BLCA cohort. Mechanistic studies incorporated RNA sequencing and ATAC-seq following histone deacetylase (HDAC) inhibition. Functional dependencies were assessed using patient-derived xenograft organoids and cell line models. Clinical relevance was explored in a biomarker-enriched investigator-initiated trial. ResultsApproximately half of urothelial carcinoma tumors exhibited BAF alterations, defining a previously unrecognized chromatin-altered molecular subtype characterized by activation of proliferative programs, loss of lineage identity, and altered metabolic signaling. This subtype was enriched for transcriptomic programs associated with HDAC inhibitor sensitivity and depleted of HDAC inhibitor resistance signatures. Mechanistically, HDAC inhibition induced widespread chromatin remodeling with reduced accessibility at AP-1 and TEAD-associated regions, and downregulation of E2F- and MYC-driven transcriptional networks. Functional studies confirmed enhanced HDAC inhibition sensitivity in ARID1A-mutated cell lines and a patient-derived organoid model. Early clinical observations demonstrated a durable responder treated with HDAC inhibitors and immunotherapy. ConclusionsBAF alterations define a chromatin-dependent tumor state in urothelial carcinoma that is selectively vulnerable to HDAC inhibition. Integrating genomic, epigenomic, functional, and early clinical evidence, these findings provide a rationale for biomarker-enriched clinical trials and HDAC inhibitor-based combination strategies in urothelial carcinoma.
Bastian, W.; Meisel, J. L.; Lee, J.-H.; Shaker, N.; Griffiths, L.; Aiello, M.; Buchwald, Z.; Liu, Y.; Thompson, E. A.; Li, Z.; Douglass, E. F.; Li, X.
Show abstract
Triple-Negative Breast Cancer (TNBC) presents a significant clinical challenge due to its heterogeneity and lack of targeted treatment options, with chemotherapy and immunotherapy combinations currently serving as the main therapeutic strategy. Efforts to address TNBC heterogeneity have largely focused on classifying intrinsic cancer subtypes based on differential tumor mRNA expression, a strategy that has proven effective in hormone receptor-positive breast cancers but has yet to yield a clinically useful predictor of survival or treatment response in TNBC. We hypothesize that both the intrinsic characteristics of TNBC and the surrounding immune microenvironment influence treatment outcomes and that immune cell infiltration affects TNBC subtype classification and response variability. To explore this hypothesis, we compared the predictive and prognostic capabilities of cancer subtype-based (TNBC-type) gene signatures and immune cell deconvolution methods (CIBERSORT) within the same TNBC datasets. We found that immune cell abundance outperformed TNBC subtype-signatures and multicellular immune cell aggregates showed the highest performance of all. More specifically, aggregate immune cells associated with tertiary lymphoid structures and tumor associated macrophages/monocytes demonstrated statistically significant predictive value. These findings were confirmed in an independent cohort of 67 TNBC patients treated with neoadjuvant chemotherapy. Further, single-cell RNA sequencing analysis revealed that the predictive power of cancer-subtype could be partially explained by immune- and stromal features. Examination of single-cell resolution spatial transcriptomic data confirmed presence of TLS-like, TAM- and cancer-stromal niches within TNBC biopsy samples that were associated with treatment response. Overall, our results highlight that immune cell aggregates, which capture the spatial organization of the TME, outperform cell-type specific gene signatures in predicting TNBC outcomes. Our novel approach provides a robust framework for interpreting spatial relationships in bulk RNA-seq data, offering a pathway for reconciling past data with current advancements in spatial profiling technologies. This work paves the way for future studies to leverage the multi-cellular complexity of TNBC, enhancing diagnostic precision and facilitating the development of therapies that strategically modulate the tumor microenvironment for improved anti-cancer responses.
JASIM, S. M.; Hezil, N.; Bouridane, A.; Hamoudi, R.
Show abstract
Accurate prognosis in lung adenocarcinoma (LUAD) requires integration of high-dimensional transcriptomic profiles with compact but clinically stable patient covariates. Naive fusion strategies allow the high-variance RNA-seq modality to dominate learned representations, suppressing clinical signal. We present Cooperative Modular Representation Learning (CMRL), an uncertainty-gated multimodal framework that dynamically regulates inter-modality information flow based on sample-level epistemic uncertainty estimated via Evidential Deep Learning (EDL). Each modality encoder produces a latent embedding and a scalar uncertainty score; an adaptive communication gate controls how much each module updates its representation from messages sent by the other module. A Variational Information Bottleneck (VIB) on the transcriptomic encoder further suppresses noise in the high-dimensional genomic latent space. CMRL is evaluated via 5-fold stratified cross validation on 490 TCGA-LUAD patients with matched RNA-seq (504 features) and clinical data. It achieves a concordance index (C-index) of 0.732 {+/-} 0.024, AUROC of 0.772 {+/-} 0.019, and AUPRC of 0.773 {+/-} 0.056 for 3-year survival prediction, outperforming a concatenation-fusion baseline (C-index 0.656), RNA-only (0.711), and clinical-only (0.670) variants, as well as several published LUAD survival models including CustOmics (0.625) and a whole-slide imaging method (0.675). An ablation study confirms that the uncertainty gate and evidential heads each contribute independently to the gain. Calibration analysis yields an Expected Calibration Error of 0.122, and uncertainty-stratified evaluation shows that low-uncertainty patients achieve AUROC 0.795 versus 0.681 for high-uncertainty patients, providing interpretable evidence that the gate mechanism is functioning as intended.
Arif, A.; Filho, J. V. d. S.
Show abstract
The increasing use of tumor sequencing has intensified the need for fast, traceable interpretation of genomic variants. General-purpose large language models can produce fluent answers, but unsupported statements, weak provenance, and stale knowledge limit their suitability for clinical genomics. We developed OncoGenRAG, a research framework that combines a parameter-efficiently fine-tuned BioBERT classifier with an entity-aware retrieval system over a curated, multi-source oncology knowledge base. The reported knowledge base contains 933 harmonized records derived from CIViC, ClinVar/dbSNP, Open Targets, UniProtKB/Swiss-Prot, Ensembl Variation, and linked PubMed literature. The classifier assigns one of five labels: Pathogenic, Likely Pathogenic, Variant of Uncertain Significance, Benign, or Oncogenic; the retrieval component ranks evidence records using subword TF-IDF similarity and explicit gene, variant, and cancer-type matches. A rejection rule suppresses answers when retrieval support is below a prespecified threshold. In the authors held-out evaluation, the classifier achieved 92.40% accuracy, 93.15% weighted precision, 92.40% weighted recall, and 92.65% weighted F1 score. In a separate benchmark of 100 clinical-style queries, OncoGenRAG achieved reported Precision@1 of 94.5%, Precision@3 of 96.8%, and 100% database grounding. No hallucinated answer was observed under the study operational definition, compared with a 41.0% no-hallucination rate for the ungrounded baseline. These results should be interpreted as internal validation rather than proof of universal safety because query construction, annotator agreement, class-specific performance, calibration, and external validation data were not available for independent analysis. OncoGenRAG provides a transparent design for evidence retrieval and abstention, but it is a research prototype and must not be used to select treatment without expert review.
Prieto-Fernandez, L.; Martinez-Carrillo, A.; de Villalain, L.; Garcia-Torre, A.; de Luxan-Delgado, B.; Hermida-Prado, F.; Navarro-Lerida, I.; Ribas, C.; Garcia-Escudero, R.; Rodrigo, J. P.; de Vicente, J. C.; Rodriguez-Santamarta, T.; Garcia-Pedrero, J. M.; Alvarez-Teijeiro, S.
Show abstract
Head and neck squamous cell carcinoma (HNSCC) remains clinically challenging, with limited molecularly targeted options and a strong dependence on the tumor microenvironment. Cancer-associated fibroblasts (CAFs) are major stromal regulators that shape tumor progression, extracellular matrix remodeling, invasion, and therapeutic response. However, how CAF heterogeneity and plasticity translate into distinct tumor-promoting functions and targetable vulnerabilities remains insufficiently defined. Here, we integrated patient-matched primary CAFs and normal fibroblasts with 3D functional assays, tumor-stroma co-culture models, quantitative extracellular matrix analysis, whole-proteome profiling, and pharmacological perturbation. Primary fibroblast populations displayed marked interpatient heterogeneity and context-dependent plasticity in invasion, contractility, and responsiveness to tumor-derived signals, whereas enhanced fibronectin-rich matrix deposition and disorganization emerged as a conserved CAF-associated feature. Both normal fibroblasts and CAFs promoted HNSCC cell invasion in a population-dependent manner, whereas CAFs consistently induced less compact and more dispersed tumor nest architectures. Integrative functional analyses identified distinct CAF phenotypes characterized by either invasive and matrix-remodeling activity or high responsiveness to tumor-derived cues. Proteomic profiling revealed recurrent enrichment of adhesion, cytoskeletal, and extracellular matrix programs and guided the selection of pharmacological inhibitors aimed at modulating specific CAF-mediated pro-tumoral functions. Pharmacological targeting selectively altered these functions: CHI3L1 inhibition disrupted fibronectin matrix deposition, broad phosphodiesterase inhibition increased matrix alignment, and FZD7 inhibition consistently blocked tumor-induced CAF invasion across all tested populations. These findings define functionally distinct and pharmacologically targetable CAF programs in HNSCC and support stromal-directed interventions as a rational component of future combination treatment strategies.
Bondarenko, M.; Qi, K.; Nowroozi, A.; Kim, J.; Kunzang, B.; Lee, A.; Liu, J.; Tran, N.; Weng, S.; Vella, M.; Chaudhari, G.; Schnizler, T.; Innanje, A.; Chen, T.; Sohn, J. H.
Show abstract
Background: Prediction of subsolid pulmonary nodule (SSN) progression from baseline CT may improve risk stratification and surveillance planning, but prior approaches have largely relied on fixed follow-up intervals. Methods: This retrospective single-center study evaluated interval-aware temporal imaging models for predicting future SSN growth and morphology across heterogeneous surveillance durations. A total of 24,946 longitudinal scan pairings derived from 2,543 clinician-reviewed SSNs in 426 patients were analyzed. A discriminative deep learning model predicted interval growth from baseline CT, segmentation masks, and interscan interval information, while a temporally conditioned generative model predicted future lesion morphology. Results: The discriminative model achieved an area under the receiver operating characteristic curve of 0.772 (95% confidence interval: 0.704-0.818), with sensitivity of 80.2% and specificity of 58.7% on the test cohort. The generative model predicted future lesion morphology with a Dice similarity coefficient of 0.706 +/-0.186. Prediction performance decreased with increasing follow-up duration, although both models generalized across intervals ranging from months to years. Conclusion: Interval-aware temporal imaging models enable the prediction of future SSN growth and morphology from baseline CT while accounting for variable surveillance intervals. These findings suggest a framework for time-aware, personalized risk assessment that may support individualized surveillance strategies and future AI-assisted management of pulmonary adenocarcinoma spectrum lesions.
Rounds, C. C.; Ravi, D.; Huang, G.; Mengesha, B.; Tran, S.; Garcia, A.; Rueb, N.; Chang, Y. H.; Park, B. S.; Wong, M. H.; Gibbs, S. L.
Show abstract
SignificanceRare-cell identification in fluorescence microscopy remains challenging because targets are sparse and background varies between specimens. Combining specimen-specific fluorescence enrichment with image classification may enable efficient and more specific automated detection of rare cells. AimWe developed a two-stage framework to identify and quantify candidate rare circulating hybrid neoplastic cells (CHCs, ECAD+/CD45+) in peripheral blood mononuclear cell (PBMC) preparations from tumor-bearing and tumor-naive mice. ApproachPBMCs from 28 mice were imaged by multichannel fluorescence microscopy. Matched unstained samples established animal-specific ECAD and CD45 background distributions for candidate cell enrichment. Blinded multi-annotator consensus labels were used to train a convolutional neural network (CNN) from DAPI, ECAD, and CD45 image crops. Generalization was evaluated by leave-one-animal-out validation across 10 random initializations. Final classification used a 10-model ensemble, and rare-cell burden was compared between groups using negative-binomial regression with total segmented-cell count as an exposure. ResultsOf the 1,065,512 segmented cells, enrichment retained 10,176 candidates (0.96%), reducing the search space by >99%. Four of five evaluable tumor-bearing animals showed reproducible held-out discrimination, with median quantified area under the receiver operator characteristic curve (AUROCs) of 0.918-0.951; one animal was a reproducible outlier (median AUROC, 0.338). Ensemble deployment identified 157.94 positive-consensus cells per 50,000 segmented cells in tumor-bearing animals versus 49.55 in controls. The estimated rare-cell rate was 3.15-fold higher in tumor-bearing animals (95% CI, 0.91-10.99; two-sided p=0.071; prespecified one-sided p=0.036). ConclusionsSpecimen-specific fluorescence enrichment combined with supervised image classification reduced the cellular search space and enabled automated quantification of a rare CHC (ECAD+/CD45+) phenotypes. Cross-animal validation also identified specimen-specific generalization failure, highlighting the importance of biological-specimen-level validation.
Chi, W. Y.
Show abstract
Background: Trophoblast cell surface antigen 2 (TROP2, encoded by TACSTD2) is a transmembrane glycoprotein overexpressed in multiple aggressive epithelial carcinomas. While antibody drug conjugates targeting TROP2 have achieved regulatory approvals, acquired payload resistance and systemic off-target toxicities limit sustained remissions. Radionuclide Drug Conjugates (RDCs) represent a potent alternative modality capable of delivering cytotoxic ionizing radiation directly to target cells. However, selecting the optimal therapeutic radioisotope between long-range beta emitters (177Lu) and short-range, high linear energy transfer (LET) alpha emitters (225Ac) under heterogeneous TROP2 spatial distributions remains an unaddressed clinical challenge. Methods: We developed an automated computational pathology and spatial microdosimetry pipeline to resolve microscopic TROP2 expression gradients and simulate absorbed radiation dose distributions from digitized whole-tissue immunohistochemistry (IHC) sections (N = 14). Optical density matrices were de-convoluted in Hematoxylin-Eosin-DAB (HED) color space to isolate the DAB chromogen. Continuous 2D spatial density distributions and topological surface profiles were reconstructed. Physical radiation energy deposition was modeled using radial dose point kernels for 177Lu (mean range ~670 m, LET 0.2 keV/m) and 225Ac (mean range ~65 m, LET 100 keV/m, 4 alpha particles per decay cascade). Therapeutic Index (TI, ratio of mean target to non-target absorbed dose), target coverage, and spatial specificity were quantified across all specimens. Results: Quantitative image deconvolution revealed that TROP2 expression across the cohort was characteristically focal and clustered, with a mean positive area fraction of 1.55 +/- 2.22% (range: 0.08% to 6.85%) and mean DAB signal intensity of 0.256 +/- 0.043. In all 14 evaluated specimens (100%), 225Ac-labeled RDCs demonstrated superior tumor-to-stroma dose localization compared to 177Lu-labeled RDCs. The cohort-wide mean Therapeutic Index was significantly higher for 225Ac (1.26 +/- 0.14) than for 177Lu (1.01 +/- 0.02, p < 0.0001, paired two-tailed t-test). Because the path length of 177Lu beta particles exceeded target cell nest dimensions by up to 30-fold, 177Lu suffered from severe off-target crossfire spillover into antigen-negative stroma. In contrast, 225Ac confined high-LET ionization tracks strictly within the micro-geographic boundaries of TROP2-expressing clusters. Conclusions: In tumors displaying focal or sparse TROP2 micro-architecture, Targeted Alpha Therapy with 225Ac-RDCs offers a superior biophysical profile over beta-emitting 177Lu-RDCs, maximizing cluster cell kill while sparing adjacent normal tissue stroma. This computational microdosimetry framework provides a practical tool to guide rational isotope pairing in RDC drug design.